id	author	title	date	pages	extension	mime	words	sentence	flesch	summary	cache	txt
bracis-33554	Reis, Willy Arthur Silva; Delgado, Karina Valdivia; Freire, Valdinei	A Unified Framework for Average Reward Criterion and Risk	2024		.htm	text/html	6092	295	61	Section 2 reviews the definitions of MDP, average reward criterion, and risk-sensitive average reward literature. 2.1 Risk-Neutral Criterion The expected total reward of a policy \(\pi \) from the initial state s up to the decision epoch \(N+1\) is a function \(v^\pi _{N+1}\) defined by $$\begin{aligned} v^\pi _{N+1}(s) = E \Bigg \{ \sum _{n=1}^{N} r(S_n,A_n) \Big | S_1=s,\pi \Bigg \} = E \Bigg \{ \sum _{n=1}^{N} r_n \Bigg \}, \end{aligned}$$ (1) where \(S_n\) and \(A_n\) refer to the random variables of the state and action in the time step n.	cache/bracis-33554.htm	txt/bracis-33554.txt
