id	author	title	date	pages	extension	mime	words	sentence	flesch	summary	cache	txt
bracis-19021	Lovatto, Ângelo Gregório; Bueno, Thiago Pereira; Barros, Leliane Nunes de	Gradient Estimation in Model-Based Reinforcement Learning: A Study on Linear Quadratic Environments	2021		.htm	text/html	5654	340	52	It is not clear, however, if better policy gradient estimation translates to more stability or faster convergence in SVG algorithms. 3.3 Stochastic Value Gradient Methods In the broader RL context, methods that learn parameterized policies, often called policy optimization methods, have gained traction in the recent decade.	cache/bracis-19021.htm	txt/bracis-19021.txt
