Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 1 https://internationalpubls.com Constructing a Precise Method to Control Non-Linear Systems Employing Special Functions and Machine Learning S.K. Sahani1*, C. Yadav2, K. Sahani3*, and V.V. Singh4 1Department of Mathematics, Janakpur Campus, T.U., Janakpurdham, Nepal *2Institute of Science and Technology, Patan Multiple Campus, T.U., Nepal 3Department of Civil Engineering, Kathmandu University, Dhulikhel, Nepal 4Professor Mathematics, Department Lingaya’s Vidhyapeeth, Faridabad, India Email:drsks53086@gmail.com1*,chandrakantyadava@gmail.com2,kameshwar.sahani@ku.edu.np3*,singh_vijayvir@yahoo.com4 Article History: Received: 08-05-2023 Revised: 02-06-2023 Accepted: 15-07-2023 Abstract: This paper is about finding the best way to control certain types of equations that describe heat and diffusion using methods that are not exact but close to the best solution. A new way to control things better is suggested using patterns we've seen before and computer learning. By employing the data gathered from the EEFs are computed using the Karhunen-Loève decompose in the PDE computer. The EEFs are used to convert the PDE system into a high-order ODE system. A smaller model (ROM) that may depict the primary behavior of the PDE subsystem is created using the SP approach. This facilitates understanding of the system. In order to reduce its size, a more basic model (ROM) that demonstrates the primary motions of the PDE system is created using the SP approach. Next, the optimal controller for the ROM is designed using the HJB technique. This ensures that the high-order ODE system is stable using the SP theory. By splitting the best way to control something into two sections, we can solve a math problem to figure out the first part and come up with a new equation for the second part. Third, a method is suggested for updating and controlling based on small improvements to solve a specific equation, and it has been proven to work effectively. Moreover, we use a technique called the NN approach to estimate the cost function. Finally, the simulation results demonstrate the effectiveness of the proposed optimum control approach when applied to a diffusion-reaction process using a unique spatially component. Keywords: The HJB equation, KLD, NN, PDE systems, optimal control, and SP are all complicated mathematical concepts. 1. Introduction Optimal control theory is a useful method for making controllers. There have been numerous significant discoveries regarding the optimal methods for regulating linear or nonlinear systems characterized by ordinary differential equations. These elements include the linear parabolic controller hypothesis, Bellman dynamic programming, and the application of the greatest assumption. [1], [2]. In real life, many industrial processes are spread out in different places, so how they work depends on where they are as well as when they are. These systems are often explained using a group of complex math equations with specific boundary rules. - Optimal control methods are difficult to use for designing a controller in real-time with a regular computer due to the complex dimensions of PDE systems. In the last few decades, Rutkowski [3] has done important work in developing the best way to control PDE systems [3], Wang, Tung, [4], and Lions [5] looked at math stuff, and you can find more theoretical results in other places too. Meanwhile, Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 2 https://internationalpubls.com people have studied how to control PDE systems effectively. Engineers have worked on ways to minimize performance measures using [6] LQ methods [7]. Meanwhile, people have studied how to control PDE systems from an engineering perspective. They have focused on methods that use LQ performance indices to minimize and improve control. Previous studies on how to make the best controller for PDE systems [8]–[10], can be divided into two types: first designing and then simplifying, and first simplifying and then designing approaches. In the first, optimal controller for PDE networks are designed using sophisticated mathematics. [8] These controllers are then simplified for use. Curtain, Zwart, Aksikas and their colleagues [9] Using algebraic operator Riccati formulae, I was able to solve the LQ ideal control issue for devices with a lot of states. Aksikas together with others. [10] I used a method called spectral factorization to find the feedback operator by solving a special type of mathematical equation. A novel approach to linear parabolic PDE system control was developed by Ray at the age of [11]. He used the original PDE model and found a way to control the system by figuring out a unique equation. Conversely, nevertheless, the reduce-then-design methods first turn the PDE system into a simpler ODE model and then use it for control design [12]–[19]. A way to control the amount of chemicals in a reactor was suggested in a book. [12] Also, two methods to control chemical reactions were created using different techniques. The investigation was undertaken by Xu and his peers. [13] They studied a type of math problem called a bilinear parabolic PDE system. A plan to make a not-perfect controller using a simple model. [14] Sadek and Bokhari used special math to estimate the state variables in a control problem with linear PDEs. [15] The Navier-Stokes equation was outfitted with a top-notch controller by 16-year-old Ravindra, who employed the Newton method to achieve optimal functionality. In Yadav et al. 's research, they used a method called approximate dynamic programming (ADP). [17] In the study by Padhi and others, make better microcontrollers by putting them together, [18], [19]. It should be emphasized that In the PDE problems covered in those works, the spatial differential operators are usually linear. [8]-[10], [12]-[19]. However, SDOs in actual industries are frequently nonlinear. For instance, in the process of rapid chemical vapour deposition, [11], [20], [21] Dispersion and react are nonlinear processes that rely on temperatures and conductivity of heat. [22], [23]. In this case, we want to create the best ways to control non-linear problems. When the spatial operator is not straight, it's hard or impossible to figure out the basic functions using math. To get around this problem, some ways of using data to find patterns have been used. Examples of such methods are Karhunen-Loève decomposition (KLD) [14], [16]–[25] and singular-value decomposition. [26], [27]. By making use of EEFs, it is possible to generate a smaller ODE system that is then applied in the development of a controller. However, only a few studies focused on finding the best way to control something, like Armaou and Christof ides [22]. Decreased gradient methods were used to solve the dynamical nonlinear code issue with equality limitations that resulted from the transformation of the optimum control problem. In recent years, the HJB method has been useful for designing the best way to control nonlinear systems. Some important progress has also been made for LPSs. nevertheless up until now, solving the HJB formulas with algebra or numbers has proven to be extremely difficult. Recently, researchers have used ADP to solve this problem in a specific way. [16]–[19], [28]–[36] This method has been used to Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 3 https://internationalpubls.com solve this problem in a careful and quiet way. ADP was suggested in [28]. There are two primary categories: bidirectional robust computer and heuristic dynamic programming [29]–[33], [31], [34]. More algorithms and a scheme for all ADP kinds were developed by Prokhorov and Wunsch. Check out the special issue [36] for further details on the use of ADP to provide input regulation. A method for solving the HJB problem for continuously running systems was created by Saridis and Lee, but they didn't figure out how to find the cost function for each step of the process. Beard and others [38], [39] proposed a way to solve the HJB equation by solving a series of simpler equations using a method called successive approximation. This method uses Galerkin approximation to solve the equations. Abu-Khalaf and Clark [40] and [41] created near-optimal state feedback controllers for continuous nonlinear systems where control restrictions are present, drawing inspiration from the work of [37]– [39]. Cheng et al. [42] and [43] employed neural networks instead of a policy iteration to answer the evolving HJB equation. They also applied this approach to the situation where inputs are restricted. The authors have limited knowledge of research on the most effective approach to managing PDE problems employing HJB equations and a regressive SDO. In this paper, we create a simpler controller for systems with certain equations using a special framework. We base our work on certain functions and neural networks. We start by using the KLD to figure out the EEFs using snapshots. The PDE network is constructed using the EEFs as a high-dimensional, singularly perturbed mathematical representation of ODEs. The SP technique helps to create a simpler model that focuses on the most important parts of the PDE system. The optimal controls for the system are designed using the ROM. This guarantees the long-term stability of the high-order ODE system. To solve the HJB-like equation, we suggest a method that uses gradual improvements and neural networks. We also show that this method will always work. Finally, we did some tests on a computer to see how well our suggested way of controlling a chemical reaction works. We found that it worked well. To sum up, this study's key contributions can be categorized into three main areas. Create a way to make a complex math method simpler for parabolic equations with a nonlinear part. Suggest a better way to control non-linear PDE systems and come up with a new type of equation. The SP technique proves that the closed-loop high-order ODE system is stable in the long run. Create a plan to update controls to solve a specific equation using a method that makes small changes and neural networks. Then show that this plan will work. This is how the remainder of the paper is structured: The definition of elliptic PDE systems is given in Section II. In the third part, the KLD approach and the SP technique are used to generate a ROM. The best controller is made and explained in Section IV. Section V is about testing a model, and Section VI is the ending. In short, this study has three main contributions. 1. Simplify a way to make parabolic PDE systems with a nonlinear SDO smaller using KLD and SP technique. 2. Suggest a good way to control systems that follow a certain mathematical rule, and come up with a new type of mathematical equation to help with this. The long-term stability of the closed-loop high-order ODE system is demonstrated by the SP approach. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 4 https://internationalpubls.com 3. Create a plan to update the control in order to solve a specific equation using a method where we keep improving our solution step by step and also use a neural network. Then show that this plan works and our solution keeps getting better. Symbols: Rn×m stands for all real n × m matrices as R represents n-dimensional space, and R is actual amounts. Rn and d stand for the usual way of multiplying things together and measuring their length in Rn, in that order. If a matrix M is symmetrical, M > (<) 0 means it is positive (negative) definite. The symbols Σ() and σ() represent the biggest and smallest singular values of a matrix. The small letter T written above a matrix means to switch the rows and columns. L2([z, z], Rn) is a big space of vectors with a certain kind of math inside, and it's related to a type of function called square integrable. It's used for working with vectors in a certain way, and has a specific way of measuring their size and how they relate to each other. Explanation of Curved PDE Systems In this paper, we study a group of curved PDE equations in one direction with a certain way of representing the state. yt(z, t) = A (y(z, t)) + f (y(z, t)) + B(z)u(t) (1) Dependent on the limits or rules. and the initial condition y (z, 0) = y0(z) (3) The suffixes z and t are part derivatives with a relationship with z and t, while y(z, t) is the state function with components y1(z, t) to yn(z, t). What is written as y(z, t) is the state. The letters z and t represent partial derivatives. The definition is related to the time coordinate, which is represented by the symbol [0,). The letter "u" of time "t" Rp is the input that is controlled. (y) is a curved SDO that involves some math stuff and follows a certain rule. t [0,) is the time coordinate and represents the values greater than or equal to 0. The manipulated input is u(t) Rp. The equation (y) is a curved line with some rules that involves different types of math and meets a certain condition. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 5 https://internationalpubls.com Where a1, a2, and a3 are real numbers that are positive. F (y) is a function that doesn't follow a straight line and meets the condition f (0) = 0. It is also continuous in a small neighborhood around each point. B (z) is a smooth matrix function that shows how control actions are spread out in different areas. Where genuine negative integers a1, a2, and a3 are involved. f (y) is a globally a Lipschitz continuously nonlinear vector function that satisfies f (0) = 0. The adequate smooth matrix function of dimension suitable, The distribution of control operations in spatial domains is believed to be characterized by B(z). The starting point is y0(z), continuous vector d1 and d2 are constant matrix M1, N1, M 2, and N2. In this research, it is assumed that the free- and closed-loop cases of the PDE system (1)–(3) are well-stated. Complex calculations will be needed for the well-posedness analysis of the PDE system (1)–(3), even for basic PDE networking with linear SDOs. [44]. We leave it for future research because the controller's architecture is the primary topic of this paper. The operator A has a specific domain, and the input operator is defined as u=B(z)u(t). Therefore, the nonlinear PDE system (1)-(3) can be restated in a simpler way. III means the third in a sequence or the roman numeral for 3. Simplifying the Reduced-Order ODE System Formula The nonlinear SDO of PDE structure (5) prevents the computation of the mathematical formulae for the eigenvalues and eigenfunctions of the PDE system. As a result, they prohibit the direct creation of finite-dimensional approximations of the PDE system using conventional basis symbol sets in combination with model reduction strategies like Galerkin's approach and symmetric collocation. To circumvent this issue, we calculate a collection of EEFs (dominant spatial patterns) of the PDE system via the retrospective technique first, using KLD [24]. These EEFs will then be applied in conjunction with the SP method to obtain a ROM that precisely explains the prevailing dynamics of the nonlinear PDE system. A. EEF Computation Using KLD In many engineering domains, KLD is a well-liked statistical pattern analysis technique for identifying the dominating structure in a collection of a multidimensional process and deriving low-dimensional accurate descriptions. Given a set of data, KLD gives an estimate of how much each EEF contributes in relation to the group's total "electricity" (mean-square fluctuation). In addition, it generates a collection of orthogonal EEFs for the mathematical model that comprise the community. The image of the trimmed series of games, as has a lower mean-square error than the other basis of one dimensions is thus best served by the EEFs as an ideal basis. Stated differently, more energy is captured by the Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 6 https://internationalpubls.com projection onto the first few EEFs than by any subsequent projection. Because of these characteristics, the EEFs are a suitable choice to take into account while reducing models. We now quickly go over the KLD process for composing EEFs using the quick snapshot technique within the context of divergent dynamic PDE circuits. Collect M sampling states of the PDE system (with u 0) (5), represented as yi(z), so that the set is sufficiently big (referred to as an ensemble). yi(z), i = 1,...,M, are referred to as "pictures" of the (5) solution. Given the assumption that the snapshots yi(z) obtained from the PDE system (5) exhibit ergodic behavior [45], the following is an expression for the spatial association value: The Mercer theorem [46] states that R(z, ξ) possesses a characteristic that where the orthogonal eigenfunctions of R(z,ξ) are represented by φi(z) while the positive eigenvalues by λi. Taking into account the subsequent integral and employing (6), we've Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 7 https://internationalpubls.com Where, αji = λ−i 1ϑj. Observe that the EEFs {\i(z)} are merely linear mixtures of snapshots {yi(z)}; hence, the computation of the coefficients is all that is needed to compute the EEFs. [α1 ··· α1i 0; Q (xs) = 0 only in that case. R is a subset of Rp×p. When ∀xs = 0, Q (xs) > 0; Q (xs) = 0 only when xs = 0. Here, Q (xs) is an optimistic continuous function. A positive definite matrix is denoted by R ∈ Rp×p. When xs = 0, V (xs) > 0; that is, V(xs) = 0 only in the event xs = 0. For xs = Ω − Rn, V (xs) − C1(Ω) is a positive definite function. We present the concept of acceptable control as follows. Criterion 1[37], [38] (acceptable Control): Given system (18), with xs ©, a control variable u(xs): Ω · Rp is determined to be acceptable with regard to (19) on Ω, represented by u(xs) (Ω), provided that the following requirements are met: First, u is continual on Ω; second, u(0) = 0; third, u(xs) stabilises the system (18); and fourth, V (xs0) <, xs0 `. Our goal is to minimise the performance functional (19) in order to identify an acceptable control u∗(xs) (Ω). It is widely acknowledged that the problems that follow may be correspondingly solved by solving the optimum control issue regarding the system in question. HJB with the Bernoulli component denoted by H(u). H (u) =Δ V ∗ T (A x + B u + g ) + Q + uTRu in which the partial derivative exists about x is indicated by the exponent xs, meaning that V ∗ =Δ ∂V ∗/∂x, and V ∗ is the magnitude of the variable. The initial-order required condition of optimality, which is satisfied by the best possible control u∗, is ∂H(u)/∂u u=u∗ = 0. As a result, we have Since u∗, the best command, violates ∂H(u)/∂u u=u∗ = 0, the situation is Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 11 https://internationalpubls.com Conversely, the related expense variable V for an uncontrolled parameter u (Ω) satisfies the following GHJB equation [37]–[39]: Substituting u∗ into (22) yields (18) Selecting V ∗ as the potential Lyapunov product for the system as a whole (18), we are left with This implies that the closed-loop system of (18) with optimal control is differently stable. u∗. The HJB equation (20) may be rewritten as follows from (21) and (23). Thus, the HJB equation (25) for V ∗ may be solved by reducing the issue of constructing an optimum control (21) to that of solving it. The two components of the dynamics of the system (18) are the negative term gs and the linear portion Asxs, as we can see. In this research, we suggest an optimum microcontroller synthesizing approach employing the following modal feedback control rule to fully use the quadratic part: Let us examine a few different versions of the monetary value V (xs) and the state penalty function Q(xs): The optimum controller may be expressed similar as (28) + (21), where Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 12 https://internationalpubls.com P and V˰∗ have yet to be ascertained. It follows from (22) that a novel kind of GHJB-like formula is provided by employing (26)–(29). [From (24)] it is evident that the synthesized ideal system for control (30) [which is equivalent to (21)] can both maximize velocity across the ROM and ensure the ROM's maximal closed-loop applications stability (18). Additionally, (12)'s closed-loop architecture is highly stable, as demonstrated by Theorem 1. In the case of the nonlinear parabolic PDE system (5), the first assumption is false is true, is the subject of Theorem 1. Next, if xs0 \1, xf0 \2, and ε (0, ε∗), thus, in an identical manner, the optimal operator (30) [or (21) may guarantee that the closed-loop system of (12) is asymmetrically stable. This is because there are positive real values δ1, δ2, and ε∗. Note that the best controller can make sure the system stays stable over time. This is because the solution obtained through a specific equation will keep getting closer to a specific value.- Additionally, [20, Proposition 1] shows that the PDE scheme (5)'s resolution y(z, t) will approach the solution yˆ(z, t) as M grows. Actually, because of the curved shape of the SDO, we can accurately describe the main way parabolic PDE systems behave using just a few patterns. Several studies have shown that a very small M can give good results, meaning that yˆ(z, t) and y(z, t) behave similarly. The top control system for ROM (18) can achieve results that are almost as good as the original PDE system (5). V. Conclusion Within the context of HJB theories, an approximation optimum control system has been synthesized in this research for a group of parabolic PDE solutions that are nonlinear. The prevailing dynamics of the PDE system may be accurately characterised by the ROM that is created using the data-based KLD method and SP technique, because the SDO of the PDE component is nonlinear. The high-order ODE system's closed-loop systems terminal stability is ensured, and an estimated optimum controller is offered by two sections to fully use the linear portion of the ROM. The linear portion of the ROM, which is acquired by solving and ARE, is the target of the first section of the synthesized controller. Conversely, the latter portion is which NN is used in the cost function estimation process. Ultimately, the case the investigation's stochastic diffusion–reaction processes is used to demonstrate the effectiveness of the mechanism of control that was developed, according to the simulation findings. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 13 https://internationalpubls.com Refrences [1] F. L. Lewis and V. L. Syrmos, Optimal Control. New York: Wiley, 1995. [2] D. P. Bertsekas, Dynamic Programming and Optimal Control, 3rd ed. Nashua, NH: Athena Scientific, 2005. [3] A. G. Butkovskiy, Distributed Control Systems. New York: American Elsevier, 1969. [4] P. K. C. Wang and F. Tung, “Optimum control of distributed parameter systems,” J. Basic Eng., vol. 86, no. 1, pp. 67–79, Mar. 1964. [5] J. L. Lions, Optimal Control of Systems Governed by Partial Differential Equations. New York: Springer-Verlag, 1971. [6] H. O. Fattorini, Infinite Dimensional Optimization Theory and Optimal Control. Cambridge, U.K.: Cambridge Univ. Press, 1999. [7] X. Li and J. Yong, Optimal Control Theory for Infinite Dimensional Systems. Boston, MA: Birkhäuser, 1995. [8] R. F. Curtain and H. J. Zwart, An Introduction to Infinite-Dimensional Linear Systems Theory. New York: Springer-Verlag, 1995. [9] I. Aksikas, A. Fuxman, J. F. Forbes, and J. J. Winkin, “LQ control design of a class of hyperbolic PDE systems: Application to fixed-bed reactor,” Automatica, vol. 45, no. 6, pp. 1542–1548, Jun. 2009. [10] I. Aksikas, J. J. Winkin, and D. Dochain, “Optimal LQ-feedback regula- tion of a nonisothermal plug flow reactor model by spectral factorization,” IEEE Trans. Autom. Control, vol. 52, no. 7, pp. 1179–1193, Jul. 2007. [11] W. H. Ray, Advanced Process Control. New York: McGraw-Hill, 1981. [12] M. Li and P. D. Christofides, “An input/output approach to the optimal transition control of a class of distributed chemical reactors,” Chem. Eng. Sci., vol. 62, no. 11, pp. 2979–2988, 2007. [13] M. Li and P. D. Christofides, “Optimal control of diffusion–convection– reaction processes using reduced-order models,” Comput. Chem. Eng., vol. 32, no. 9, pp. 2123–2135, Sep. 2008. [14] C. Xu, Y. Ou, and E. Schuster, “Sequential linear quadratic control of bilinear parabolic PDEs based on POD model reduction,” Automatica, vol. 47, no. 2, pp. 418–426, Feb. 2011. [15] I. S. Sadek and M. A. Bokhari, “Optimal control of a parabolic distrib- uted parameter system via orthogonal polynomials,” Optim. Control Appl. Methods, vol. 19, no. 3, pp. 205–213, May/Jun. 1998. [16] S. S. Ravindran, “A reduced-order approach for optimal control of fluids using proper orthogonal decomposition,” Int. J. Numer. Method Fluid, vol. 34, no. 5, pp. 425–448, Nov. 2000. [17] V. Yadav, R. Padhi, and S. N. Balakrishnan, “Robust/optimal temperature profile control of a high-speed aerospace vehicle using neural networks,” IEEE Trans. Neural Netw., vol. 18, no. 4, pp. 1115–1128, Jul. 2007. [18] R. Padhi, S. N. Balakrishnan, and T. Randolph, “Adaptive-critic based optimal neuro control synthesis for distributed parameter systems,” Automatica, vol. 37, no. 8, pp. 1223–1234, 2001. [19] R. Padhi and S. N. Balakrishnan, “Proper orthogonal decomposition based optimal neurocontrol synthesis of a chemical reactor process using approximate dynamic programming,” J. Neural Netw., vol. 16, no. 5/6, pp. 719–728, Jun. 2003. Communications on Applied Nonlinear Analysis ISSN: 1074-133X Vol 30 No. 2 (2023) 14 https://internationalpubls.com [20] J. Baker and P. D. Christofides, “Finite-dimensional approximation and control of nonlinear parabolic PDE systems,” Int. J. Control, vol. 73, no. 5, pp. 439–456, 2000. [21] J. Baker and P. D. Christofides, “Output feedback control of parabolic PDE systems with nonlinear spatial differential operators,” Ind. Eng. Chem. Res., vol. 38, no. 11, pp. 4372–4380, Nov. 1999. [22] A. Armaou and P. D. Christofides, “Dynamic optimization of dissipative PDE systems using nonlinear order reduction,” Chem. Eng. Sci., vol. 57, no. 24, pp. 5083–5114, Dec. 2002. [23] H.-X. Li and C. K. Qi, Spatio-Temporal Modeling of Nonlinear Dis- tributed Parameter Systems: A Time/Space Separation Based Approach. New York: Springer-Verlag, 2011. [24] P. Holmes, J. L. Lumley, and G. Berkooz, Turbulence, Coherent Struc- tures, Dynamical Systems and Symmetry. Cambridge, U.K.: Cambridge Univ. Press, 1996. [25] L. G. Bleris and M. V. Kothare, “Low-order empirical modeling of dis- tributed parameter systems using temporal and spatial eigenfunctions,” Comput. Chem. Eng., vol. 29, no. 4, pp. 817–827, Mar. 2005. [26] D. H. Gay and W. H. Ray, “Identification and control of distributed parameter systems by means of the singular value decomposition,” Chem. Eng. Sci., vol. 50, no. 10, pp. 1519–1539, May 1995. [27] D. G. Zheng, “System identification and model-based control for a class of distributed parameter systems,” Ph.D. dissertation, Texas Tech Univ., Lubbock, TX, 2003. [28] B. Widrow, N. K. Gupta, and S. Maitra, “Punish/reward: Learning with a critic in adaptive threshold systems,” IEEE Trans. Syst., Man, Cybern., vol. SMC-3, no. 5, pp. 455–465, Sep. 1973. [29] J. Seiffertt, S. Sanyal, and D. C. Wunsch, “Hamilton–Jacobi–Bellman equations and approximate dynamic programming on time scales,” IEEE Trans. Syst., Man, Cybern. B, Cybern., vol. 38, no. 4, pp. 918–923, Aug. 2008. [30] H. Zhang, Q. Wei, and Y. Luo, “A novel infinite-time optimal tracking control scheme for a class of discrete-time nonlinear systems via the greedy HDP iteration algorithm,” IEEE Trans. Syst., Man, Cybern. B, Cybern., vol. 38, no. 4, pp. 937–942, Aug. 2008. [31] A. Al-Tamimi, M. Abu-Khalaf, and F. L. Lewis, “Adaptive critic designs for discrete-time zero-sum games with application to H control,” IEEE Trans. Syst., Man, Cybern. B, Cybern., vol. 37, no. 1, pp. 240–247, Feb. 2007. [32] A. Al-Tamimi, F. L. Lewis, and M. Abu-Khalaf, “Discrete-time nonlin- ear HJB solution using approximate dynamic programming convergence proof,” IEEE Trans. Syst., Man, Cybern. B, Cybern., vol. 38, no. 4, pp. 943–949, Aug. 2008. [33] F. L. Lewis and K. G. Vamvoudakis, “Reinforcement learning for partially observable dynamic processes: Adaptive dynamic programming using measured output data,” IEEE Trans. Syst., Man, Cybern. B, Cybern., vol. 41, no. 1, pp. 14–25, Feb. 2011. [34] H. Zhang and Y. Luo, “Neural-network-based near-optimal control for a class of discrete-time affine nonlinear systems with control constraints,” IEEE Trans. Neural Netw., vol. 20, no. 9, pp. 1490–1503, Sep. 2009.