id	author	title	date	pages	extension	mime	words	sentence	flesch	summary	cache	txt
dad-11518	Ultes, Stefan; Maier, Wolfgang	User Satisfaction Reward Estimation Across Domains: Domain-independent Dialogue Policy Learning	2021	34	.pdf	application/pdf	15869	764	4	Relevant Related Work Most of previous work on dialogue policy learning focuses on employing task success as the main reward signal (Gašić and Young, 2014; Gašić et al., 2014; Lemon and Pietquin, 2007; Daubigney et al., 2012; Levin and Pieraccini, 1997; Singh et al., 2002; Young et al., 2013; Su et al., 2015, 2016). This would mean that it would be less suitable to be applied to dialogue policy learning for different domains.	cache/dad-11518.pdf	txt/dad-11518.txt
